NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

An Empirical Evaluation of Low-Rank Adapted Vision–Language Models for Radiology Image Captioning

https://doi.org/10.3390/bioengineering12121330

Hoque, Mahmudul; Chowdhury, Raisa Nusrat; Hasan, Md Rakibul; Peter, Ojonugwa_Oluwafemi Ejiga; Khalifa, Fahmi; Rahman, Md Mahmudur (December 2025, Bioengineering)

Rapidly growing medical imaging volumes have increased radiologist workloads, creating demand for automated tools that support interpretation and reduce reporting delays. Vision-language models (VLMs) can generate clinically relevant captions to accelerate report drafting, yet their varying parameter scales require systematic evaluation for clinical utility. This study evaluated ten multimodal models fine-tuned on the Radiology Objects in Context version 2 (ROCOv2) dataset containing 116,635 images across eight modalities. We compared four Large VLMs (LVLMs) including LLaVA variants and IDEFICS-9B against four Small VLMs (SVLMs) including MoonDream2, Qwen variants, and SmolVLM, alongside two fully fine-tuned baseline architectures (VisionGPT2 and CNN-Transformer). Low-Rank Adaptation (LoRA), applied to fewer than 1% of selected model parameters, proved optimal among adaptation strategies, outperforming broader LoRA configurations. Models were assessed on relevance (semantic similarity) and factuality (concept-level correctness) metrics. Performance showed clear stratification: LVLMs (0.273 to 0.317 overall), SVLMs (0.188 to 0.279), and baselines (0.154 to 0.177). LLaVA-Mistral-7B achieved the highest performance with relevance and factuality scores of 0.516 and 0.118, respectively, substantially exceeding the VisionGPT2 baseline (0.325, 0.028). Among the SVLMs, MoonDream2 demonstrated competitive relevance (0.466), approaching the performance of some LVLMs despite its smaller size. To investigate performance enhancement strategies for underperforming SVLMs, we prepended predicted imaging modality labels at inference time, which yielded variable results. These findings provide quantitative benchmarks for VLM selection in medical imaging, demonstrating that while model scale influences performance, architectural design and targeted adaptation enable select compact models to achieve competitive results.
more » « less
Full Text Available
Fair Federated Survival Analysis

https://doi.org/10.1609/aaai.v39i19.34214

Rahman, Md Mahmudur; Purushotham, Sanjay (April 2025, Proceedings of the AAAI Conference on Artificial Intelligence)

Federated Survival Analysis (FSA) is an emerging Federated Learning (FL) paradigm that enables training survival models on decentralized data while preserving privacy. However, existing FSA approaches largely overlook the potential risk of bias in predictions arising from demographic and censoring disparities across clients' datasets, which impacts the fairness and performance of federated survival models, especially for underrepresented groups. To address this gap, we introduce FairFSA, a novel FSA framework that adapts existing fair survival models to the federated setting. FairFSA jointly trains survival models using distributionally robust optimization, penalizing worst-case errors across subpopulations that exceed a specified probability threshold. Partially observed survival outcomes in clients are reconstructed with federated pseudo values (FPV) before model training to address censoring. Furthermore, we design a weight aggregation strategy by enhancing the FedAvg algorithm with a fairness-aware concordance index-based aggregation method to foster equitable performance distribution across clients. To the best of our knowledge, this is the first work to study and integrate fairness into Federated Survival Analysis. Comprehensive experiments on distributed non-IID datasets demonstrate FairFSA's superiority in fairness and accuracy over state-of-the-art FSA methods, establishing it as a robust FSA approach capable of handling censoring while providing equitable and accurate survival predictions for all subjects.
more » « less
Full Text Available
Parameter-Efficient VLMs for Gastrointestinal Endoscopy: Medical Image Generation and Clinical Visual Question Answering

https://doi.org/10.1109/BHI67747.2025.11269532

Ejiga_Peter, Ojonugwa Oluwafemi; Akor_Ejiga, Frederick; Khalifa, Fahmi; Rahman, Md Mahmudur (October 2025, IEEE)

Full Text Available
Text-Guided Synthesis in Medical Multimedia Retrieval: A Framework for Enhanced Colonoscopy Image Classification and Segmentation

https://doi.org/10.3390/a18030155

Ejiga_Peter, Ojonugwa Oluwafemi; Adeniran, Opeyemi Taiwo; John-Otumu, Adetokunbo MacGregor; Khalifa, Fahmi; Rahman, Md Mahmudur (March 2025, Algorithms)

The lack of extensive, varied, and thoroughly annotated datasets impedes the advancement of artificial intelligence (AI) for medical applications, especially colorectal cancer detection. Models trained with limited diversity often display biases, especially when utilized on disadvantaged groups. Generative models (e.g., DALL-E 2, Vector-Quantized Generative Adversarial Network (VQ-GAN)) have been used to generate images but not colonoscopy data for intelligent data augmentation. This study developed an effective method for producing synthetic colonoscopy image data, which can be used to train advanced medical diagnostic models for robust colorectal cancer detection and treatment. Text-to-image synthesis was performed using fine-tuned Visual Large Language Models (LLMs). Stable Diffusion and DreamBooth Low-Rank Adaptation produce images that look authentic, with an average Inception score of 2.36 across three datasets. The validation accuracy of various classification models Big Transfer (BiT), Fixed Resolution Residual Next Generation Network (FixResNeXt), and Efficient Neural Network (EfficientNet) were 92%, 91%, and 86%, respectively. Vision Transformer (ViT) and Data-Efficient Image Transformers (DeiT) had an accuracy rate of 93%. Secondly, for the segmentation of polyps, the ground truth masks are generated using Segment Anything Model (SAM). Then, five segmentation models (U-Net, Pyramid Scene Parsing Network (PSNet), Feature Pyramid Network (FPN), Link Network (LinkNet), and Multi-scale Attention Network (MANet)) were adopted. FPN produced excellent results, with an Intersection Over Union (IoU) of 0.64, an F1 score of 0.78, a recall of 0.75, and a Dice coefficient of 0.77. This demonstrates strong performance in terms of both segmentation accuracy and overlap metrics, with particularly robust results in balanced detection capability as shown by the high F1 score and Dice coefficient. This highlights how AI-generated medical images can improve colonoscopy analysis, which is critical for early colorectal cancer detection.
more » « less
Full Text Available
Communication-Efficient Pseudo Value-Based Random Forests for Federated Survival Analysis

https://doi.org/10.1609/aaaiss.v2i1.27714

Rahman, Md Mahmudur; Purushotham, Sanjay (October 2023, AAAI Fall Symposium Series: Second Symposium on Survival Prediction: Algorithms, Challenges, and Applications (SPACA))

Full Text Available
Federated Competing Risk Analysis

https://doi.org/10.1145/3583780.3614880

Rahman, Md Mahmudur; Purushotham, Sanjay (October 2023, ACM)
Federated learning for competing risk analysis in healthcare

Rahman, Md Mahmudur; Purushotham, Sanjay (August 2023, International Workshop on Federated Learning for Distributed Data Mining)

Full Text Available
FedPseudo: Privacy-Preserving Pseudo Value-Based Deep Learning Models for Federated Survival Analysis

https://doi.org/10.1145/3580305.3599348

Rahman, Md Mahmudur; Purushotham, Sanjay (August 2023, ACM)
A Hybrid Learning-Architecture for Mental Disorder Detection Using Emotion Recognition

https://doi.org/10.1109/ACCESS.2024.3421376

Aina, Joseph; Akinniyi, Oluwatunmise; Rahman, Md Mahmudur; Odero-Marah, Valerie; Khalifa, Fahmi (January 2024, IEEE Access)

Mental illness has grown to become a prevalent and global health concern that affects individuals across various demographics. Timely detection and accurate diagnosis of mental disorders are crucial for effective treatment and support as late diagnosis could result in suicidal, harmful behaviors and ultimately death. To this end, the present study introduces a novel pipeline for the analysis of facial expressions, leveraging both the AffectNet and 2013 Facial Emotion Recognition (FER) datasets. Consequently, this research goes beyond traditional diagnostic methods by contributing a system capable of generating a comprehensive mental disorder dataset and concurrently predicting mental disorders based on facial emotional cues. Particularly, we introduce a hybrid architecture for mental disorder detection leveraging the state-of-the-art object detection algorithm, YOLOv8 to detect and classify visual cues associated with specific mental disorders. To achieve accurate predictions, an integrated learning architecture based on the fusion of Convolution Neural Networks (CNNs) and Visual Transformer (ViT) models is developed to form an ensemble classifier that predicts the presence of mental illness (e.g., depression, anxiety, and other mental disorder). The overall accuracy is improved to about 81% using the proposed ensemble technique. To ensure transparency and interpretability, we integrate techniques such as Gradient-weighted Class Activation Mapping (Grad-CAM) and saliency maps to highlight the regions in the input image that significantly contribute to the model’s predictions thus providing healthcare professionals with a clear understanding of the features influencing the system’s decisions thereby enhancing trust and more informed diagnostic process.
more » « less
Full Text Available
Comparative Analysis of Fine-Tuned Multimodal Models in Radiology Image Captioning

https://doi.org/10.1109/ICMI65310.2025.11141312

Hoque, Mahmudul; Hasan, Md Rakibul; Emon, Md_Ismail Siddiqi; Oluwafemi, Ejiga_Peter Ojonugwa; Rahman, Md Mahmudur; Khalifa, Fahmi (April 2025, IEEE)

Full Text Available

« Prev Next »

Search for: All records